All raw data used for plotting Figures and Supplementary Figures were provided in "Figures" directory.

1)Source Data.xlsx：Most raw data of Figures are provided by a single sheet in an Excel document, sheets were named with Figure number (such as Fig. 2a, Fig. 2b, Supplementary Fig. 6). Multiple data files were required for plotting Figure 5 and Supplmentary Fig. 10 , so we useded two and three sheets which were named Fig. 5-1, Fig. 5-2, and Supp Fig. 10-1, Supp Fig. 10-2 and Supp Fig. 10-3, respectively. Figure 4a-c were plottted with the same one raw data file, we named that sheet like "Fig.4,a,b,c", the same for Figure 6.

2) Fig. 2e.xls: This dataset is very big and occupys morn than 17 million lines, we provided it as a separate document.

3) *.treefile: these three tree files (Fig. 4b.treefile, Fig. 6a.treefile, Fig. 6c.treefile) were used for visualization of phylogenetic analysis.

The scripts of metagenomic analysis are placed in "Pipeline" directory. There are two main modules in the pipeline including the construction of the gene catalog and metagenome-assembled genomes. The data processes before assembly are the same for these two modules.

Text processing in the pipeline, statistical analysis and visualization were handled by scripting with R, Shell, Perl or Python languages. These scripts were placed in "Scripts" directory. All related input data for statistical analysis and visualization are in "Pre-processed_Files" directory.